balancing risk and reward
Balancing Risk and Reward: A Batched-Bandit Strategy for Automated Phased Release
Phased releases are a common strategy in the technology industry for gradually releasing new products or updates through a sequence of A/B tests in which the number of treated units gradually grows until full deployment or deprecation. Performing phased releases in a principled way requires selecting the proportion of units assigned to the new release in a way that balances the risk of an adverse effect with the need to iterate and learn from the experiment rapidly. In this paper, we formalize this problem and propose an algorithm that automatically determines the release percentage at each stage in the schedule, balancing the need to control risk while maximizing ramp-up speed. Our framework models the challenge as a constrained batched bandit problem that ensures that our pre-specified experimental budget is not depleted with high probability. Our proposed algorithm leverages an adaptive Bayesian approach in which the maximal number of units assigned to the treatment is determined by the posterior distribution, ensuring that the probability of depleting the remaining budget is low. Notably, our approach analytically solves the ramp sizes by inverting probability bounds, eliminating the need for challenging rare-event Monte Carlo simulation. It only requires computing means and variances of outcome subsets, making it highly efficient and parallelizable.
Balancing Risk and Reward: A Batched-Bandit Strategy for Automated Phased Release
Phased releases are a common strategy in the technology industry for gradually releasing new products or updates through a sequence of A/B tests in which the number of treated units gradually grows until full deployment or deprecation. Performing phased releases in a principled way requires selecting the proportion of units assigned to the new release in a way that balances the risk of an adverse effect with the need to iterate and learn from the experiment rapidly. In this paper, we formalize this problem and propose an algorithm that automatically determines the release percentage at each stage in the schedule, balancing the need to control risk while maximizing ramp-up speed. Our framework models the challenge as a constrained batched bandit problem that ensures that our pre-specified experimental budget is not depleted with high probability. Our proposed algorithm leverages an adaptive Bayesian approach in which the maximal number of units assigned to the treatment is determined by the posterior distribution, ensuring that the probability of depleting the remaining budget is low.
AI and financial processes: Balancing risk and reward
All the sessions from Transform 2021 are available on-demand now. Of all the enterprise functions influenced by AI these days, perhaps none is more consequential than AI and financial processes. People don't like when other people fiddle with their money, let alone an emotionless robot. But as it usually goes with first impressions, AI is winning converts in monetary circles, in no small part due to its ability to drive out inefficiencies and capitalize on hidden opportunities – basically creating more wealth out of existing wealth. One of the ways it does this is to reduce the cost of accuracy, says Sanjay Vyas, CTO of Planful, a developer of cloud-based financial planning platforms.